The Dynamic Replication Mechanism of HDFS Hot File based on Cloud Storage
نویسندگان
چکیده
As an open source cloud storage scheme, HDFS is used by more and more large enterprises and researchers, and is actually applied to many cloud computing systems to deal with huge amounts of data. HDFS has many advantages, but there are some problems such as NameNode single point of failure, small file problem, hot issues, etc. For HDFS hot issues, this paper proposes a dynamic Replication mechanism of HDFS hot file based on cloud storage(HDFS-DRM). The mechanism includes a Replication of the dynamic adjustment mechanism and adding, deleting duplicate node selection mechanism in two parts, by increasing the NameNode, BlockMap parameters, it records the number of reading requests of each file in a certain period of time to decide whether to increase or decrease the number of copies. The mechanism presents a replica placement method based on stage historical information and node load and selects the appropriate node to add or delete copies of documents to improve the utilization efficiency of the data node storage space effectively. Experimental results show that, HDFS DRM in hot files case, compared to native HDFS file system access latency is significantly reduced, HDFS-DRM can solve the hot issues successfully.
منابع مشابه
An Efficient Data Replication Strategy in Large-Scale Data Grid Environments Based on Availability and Popularity
The data grid technology, which uses the scale of the Internet to solve storage limitation for the huge amount of data, has become one of the hot research topics. Recently, data replication strategies have been widely employed in distributed environment to copy frequently accessed data in suitable sites. The primary purposes are shortening distance of file transmission and achieving files from ...
متن کاملImproving Data Grids Performance by Using Modified Dynamic Hierarchical Replication Strategy
Abstract: A Data Grid connects a collection of geographically distributed computational and storage resources that enables users to share data and other resources. Data replication, a technique much discussed by Data Grid researchers in recent years creates multiple copies of file and places them in various locations to shorten file access times. In this paper, a dynamic data replication strate...
متن کاملManagement of Data Replication for PC Cluster-based Cloud Storage System
Storage systems are essential building blocks for cloud computing infrastructures. Although high performance storage servers are the ultimate solution for cloud storage, the implementation of inexpensive storage system remains an open issue. To address this problem, the efficient cloud storage system is implemented with inexpensive and commodity computer nodes that are organized into PC cluster...
متن کاملResearch and implementation on cloud computing security based on HDFS
This paper focuses on the research of the cloud computing security, proposing the file data management model and implementing the security of the cloud computing based on HDFS. The design of the file data management system under the cloud computer environment is achieved based on HDFS, which is with the functions of upload and download data parallelism, user management, inventory management, et...
متن کاملImproving Data Availability Using Combined Replication Strategy in Cloud Environment
As grow as the data-intensive applications in cloud computing day after day, data popularity in this environment becomes critical and important. Hence to improve data availability and efficient accesses to popular data, replication algorithms are now widely used in distributed systems. However, most of them only replicate the static number of replicas on some requested chosen sites and it is ob...
متن کامل